IEEE Access
● Institute of Electrical and Electronics Engineers (IEEE)
Preprints posted in the last 90 days, ranked by how well they match IEEE Access's content profile, based on 35 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit.
Biswas, J.; Islam, M.; Bangabashi, M. M.; Akter, M.; Nishi, T. S.; Sheikh, M. K.; Mia, M. R.; Anwar, M. M.
Show abstract
Guava cultivation is considerably influenced by foliar and fruit diseases whose overlapping symptoms and environmental variability make accurate field-level diagnosis challenging. Numerous studies have been conducted to find efficient methods of diagnosing plant diseases, but most focus on image-level classification and do not include lesion localization or pixel-level segmentation of the images within a single framework of analysis. This study proposes a comprehensive framework for utilizing automated image analysis to classify guava leaf and fruit diseases at the image level, locate lesions, and segment lesions at the pixel level from multiple images of the same type of disease collected from various growing conditions. The dataset was enriched through three augmentation strategies including standard preprocessing, structured augmentation, and GAN-based synthetic image generation, expanding the effective training data to approximately 7,000 images, while a 5-fold cross-validation strategy guided model selection and final performance was assessed on a held-out test set. The experimental evaluation of multiple state-of-the-art Convolutional Neural Networks (CNNs) for the classification of guava leaf and fruit diseases indicated that the model generated using the ResNet50+DenseNet121 model fusion achieved the highest classification accuracy of 98.20%. For lesion detection and segmentation, YOLOv8-seg outperformed Mask R-CNN, achieving mAP@0.5 of 0.907 and 0.889, and mAP@0.5:0.95 of 0.783 and 0.769 for detection and segmentation, respectively, with a balanced precision-recall profile. The techniques of Explainable AI (XAI) were used to increase the transparency of this model by identifying areas in the image that are significant to the actual lesion. The framework was further designed with practical web-based deployment in mind, evaluating both lightweight and high-capacity models to balance computational efficiency against predictive accuracy. From this research, it was concluded that using model fusion, data augmentation, and segmentation-aware lesion detection would provide a solution for managing guava diseases effectively.
Gonzalez Nunez, J. G.; Sabri, S.; Kebria, P.; Crook, J.; Brattain, L.
Show abstract
This paper presents a two-stage pipeline for implicit feature engineering in time series-based physiological stress detection using electrodermal activity (EDA) signals. In the first stage, we forecast three descriptive statistics of future EDA signals over short horizons (3, 5, and 10 seconds) based on a 60-second context window. In the second stage, a lightweight linear classifier detects stress from these predicted statistics. We evaluate three forecasting architectures spanning the domain expertise spectrum: a domain-specific bidirectional long short-term memory (BiLSTM) recurrent neural network, zero-shot and fine-tuned variants of Amazon Chronos T5 time series foundation model, and the Tabular Prior-data Fitted Network (TabPFN) applied to engineered physiological features. Experiments on the publicly available Wearable Stress and Affect Detection (WESAD) dataset, comprising chest-worn multimodal physiological signals from 15 subjects under baseline and stress conditions, demonstrate that the domain-specific BiLSTM achieves the highest classification performance, with area under the receiver operating characteristic curve (AUC) values ranging from 0.913 to 0.962. TabPFN follows with AUC values of 0.853-0.869, while Chronos variants yield 0.528-0.744. Notably, models using predicted features consistently outperform those using oracle features derived from the true future signals--the theoretical upper bound--suggesting effective noise filtering through learned sequence representations. Chronos models quickly reach performance saturation regardless of training depth, highlighting challenges in tokenizing continuous physiological time series. The proposed approach advances implicit feature engineering for wearable stress monitoring by leveraging forecasting as a powerful inductive bias, thereby improving robustness and providing insights into the limitations of the foundation model for physiological signals.
Zhuang, Q.; Mou, C.; Liu, B.; Fu, M. R.; King, G. W.
Show abstract
Breast cancer survivors frequently experience upper-limb impairments, making continuous monitoring essential for effective rehabilitation. We propose REINA (Recognize-Then-Infer Wearable-to-App AI Framework), a two-stage deep-learning approach for remote monitoring of motor function during breast cancer rehabilitation using wearable-device data. Inertial measurement unit (IMU) signals from wearable devices are first used to recognize physical activities via supervised learning, followed by an activity-specific recurrent neural network (RNN) to infer corresponding electromyography (EMG) signals. REINA establishes reliable inference of neuromuscular activity from wearable IMU data, enabling real-time, cost-effective assessment of motor function recovery in real-world settings.
Ghaffarzadeh, P.; Chakraborty, D.; Aslansefat, K.; Dostan, A.; Papadopoulos, Y.
Show abstract
Ground reaction force (GRF) measurement remains largely confined to instrumented laboratories, limiting longitudinal monitoring in daily life. This article presents an edge-first wearable system for estimating vertical GRF from consumer smartwatches. Two Apple Watch Series 6 devices worn at the wrist and waist stream 12-channel inertial data at 100 Hz to an iPhone, where preprocessing, storage, and inference occur locally without cloud dependence. The proposed GRFNet-MultiScale model is a compact temporal convolutional network with four dilated residual blocks and a global context branch. Under leave-one-subject-out evaluation on 539 stance windows from 10 healthy participants, the dual-sensor system achieved a mean Pearson correlation of 0.798 with an RMSE of 257 N, while a wrist-only configuration retained 82.5% of dual-sensor correlation. Temporal attribution remained stable across validation folds and identified early-stance wrist acceleration as the dominant reproducible signal. The system is strongest for cyclic locomotion.
Mlynczak, M.; Rosol, M.; Korzeniewski, K.; Gasior, J. S.
Show abstract
Background and ObjectiveAccurately parameterizing dynamic, time-varying interactions in physiological systems is a methodological challenge, as global causal discovery methods may obscure transient, local fluctuations. This study introduces tempord, an open-source Python library designed to estimate local temporal orders and evaluate the short-term stability, directionality, and strength of causal links in non-stationary biological signals. MethodsThe algorithm estimates temporal relationships by keeping one signal stationary while iteratively shifting another one within a sliding window. To parameterize optimal inter-signal shifts (causal vector, CV), the framework utilizes linear modeling or time series distance metrics. The methodology was validated through a simulation study on synthetic bivariate signals with mathematically imposed dynamic phase delays, under both deterministic and noisy conditions. Furthermore, in-vivo capabilities were demonstrated by evaluating cardiorespiratory coupling dynamics across spontaneous and music-induced relaxation breathing states. ResultsThe simulation study demonstrated that the extracted CV trajectories precisely aligned with ground-truth temporal delays, assessed using mean absolute error and root mean square error for both noise-free and noisy synthetic data. In-vivo application demonstrated dynamic temporal stability and the detection of minor step changes during autonomic nervous system state transitions. ConclusionsThe tempord Python package bridges the gap between global causal discovery and local beat-by-beat statistical parameterization. It provides a robust "bottom-up" analytical instrument for investigating the transient mechanisms governing complex biological networks.
Oladunni, T.; Ganiyu Adewumi, F.
Show abstract
Photoplethysmography (PPG, optical measurement of cardiac blood volume changes) is the foundation of wearable cardiac monitoring, but systematically fails on dark skin due to melanin absorption. We present the Melanin Absorption Invariance (MAI) framework: a label-free method that substantially reduces cross-skin-tone bias in cardiac feature extraction by preserving topological rather than geometric signal structure. We prove two theorems: Theorem 1 bounds attractor bias to O(SNR_eff^-1) under Z-normalization; Theorem 2 reduces residual bias to O(SNR_eff^-2) via SNR-adaptive correction. Empirical validation confirms these theoretical predictions on real dark-skin PPG signals. Comprehensive empirical validation on the complete MMPD dataset (Fitzpatrick III-VI, n = 656 recordings, 33 subjects, spanning all 4 lighting conditions and 5 motion types, Samsung Galaxy mobile phone) demonstrates MAI generalization across real-world deployment conditions. Results show substantial attractor bias reduction across all skin tone groups, with largest effects for Fitzpatrick IV and VI populations most affected by current systems. This work demonstrates a theoretically grounded, label-free, skin-tone-invariant cardiac monitoring framework.
Yuan, Y.; Li, W.; Zhu, L.; Su, H.; Yu, H.; Wang, H.; Lin, G. N.
Show abstract
Freezing of gait (FoG) in Parkinson's disease is a brief but hazardous gait failure that often precedes falls. For wearable cueing or other closed-loop assistance, a detector that reacts only after FoG onset is usually too late; the more useful task is to recognize the pre-freezing transition from physiological signals. This study presents PreFoGNet, a dual time-frequency deep learning framework for early FoG prediction using plantar pressure signals. The temporal stream combines a multi-scale Inception encoder with a bidirectional Mamba module to capture both short contact-related transients and several-second gait deterioration without the quadratic cost of attention. In parallel, the frequency stream uses band-wise spectral modeling and attention-based gating to emphasize physiologically meaningful changes in the locomotion, freeze-related, and high-frequency bands. On the WearGait-PD dataset, with a 2 s prediction horizon and subject-wise evaluation, PreFoGNet achieved a sensitivity of 93.94%, a specificity of 89.76%, a G-Mean of 0.9183, and an AUC-ROC of 0.9607. It outperformed classical machine-learning and deep learning baselines, and retained usable performance under moderate noise and single-channel loss. Additional horizon analysis showed that plantar pressure contains a stable pre-freezing signature within 0-3 s before onset, with a practical prediction boundary of approximately 6-7 s. These findings suggest that time-frequency modeling of plantar pressure is a promising signal-processing route for wearable FoG early-warning systems.
Fu, J.; Zhang, S.; Huang, H. J.; Rakhshan, M.; Wen, Y.
Show abstract
Motor unit (MU) decomposition using high-density surface electromyography (HD-sEMG) has been widely used to characterize MU behavior in neurophysiology and to build neural-machine interfaces for wearable robots. Recently, many open-source software tools for MU decomposition have been made available on GitHub, which could reduce the effort of researchers in the field. However, the consistency among these open-source tools has never been studied, making researchers hesitate to use them. In this study, we collected 7 open-source software tools on GitHub and applied them to decompose MUs from an open-source HD-sEMG dataset (including 11 isometric contraction trials) to investigate the consistency among these tools. To create a comprehensive MU pool for reference, we combined all unique MUs identified by seven tools, visually inspected and removed bad MUs, and manually edited all remaining MU spike trains. Across 7 tools for 11 trials, the number of identified MUs ranges from 167 to 736. The number of valid MUs after expert inspection ranges from 29 to 210, which is 10% to 72% of the reference pool. The rate of agreement between the raw MUSTs and the manually edited MUSTs ranges from 0.86 to 0.94, and the averaged number of edits per MU to correct misalignments ranges from 14 to 39. The results show inconsistency in the implementation and procedures of each tool, which results in an inconsistent number of identified MUs and valid MUs (29 vs 210). In general, a substantial amount of effort is required to process the raw MUSTs from each tool to conduct further research analysis. This study provided a guideline for using open-source software tools for MU decomposition and indicated that it would be beneficial to develop tools to automatically edit the MUSTs.
Howlader, D.; AHMED, T.; Rahman, M. M.
Show abstract
Parkinsons Disease (PD) is a progressive neurodegenerative disorder which significantly affects motor function, daily coordination and verbal communication. Speech-based biomarkers provide a non-invasive and scalable approach to early detection, as dysphonia is one of the earliest and most consistent clinical markers of PD. The dataset used in this study is publicly available and consists of 756 voice recordings from 252 subjects (188 with PD and 64 neurologically healthy controls) with a wide range of acoustic parameters such as Mel-Frequency Cepstral Coefficients (MFCCs), energy-based parameters, and higher-order statistical derivatives. After systematic preprocessing and z-score normalisation, five machine learning classifiers were tested: K-Nearest Neighbors (KNN), Extreme Gradient Boosting (XGBoost), Random Forest (RF), Support Vector Machine (SVM), and Naive Bayes (NB) under a subject-independent, GroupKFold cross-validation protocol. The KNN classifier performed best overall with an accuracy of 92.10%, F1 score of 94.50%, and a precision rate of 98.09%, reducing the number of false positive diagnoses. To overcome the lack of interpretability of black-box predictive models, SHapley Additive exPlanations (SHAP) were used to explain the contribution of each feature to the prediction of an individual. The most diagnostically salient acoustic biomarkers were identified as features from the SHAP analysis: std delta delta log energy, the first Mel-Frequency Cepstral Coefficient, and Tunable Q-Factor Wavelet Transform (TQWT). This work introduces a machine learning framework that is both reproducible and clinically interpretable, combining high predictive accuracy with transparent, physiologically grounded decision logic.
Sharma, O.;Weidenfeld, K.;Barkan, D.;Gal, O.
Show abstract
Breast cancer cells that disseminate to distant organs can remain dormant (non-proliferative) for years before reactivating and progressing into lethal metastatic disease. Understanding the transition between dormancy and reactivation is therefore critical for early intervention and treatment. In this study, we investigate a comprehensive range of deep learning (DL) architectures to classify dormant versus proliferative breast tumor cells within a 3-dimensional growth factor reduced basement membrane extract (3D BME) system that models tumor dormancy and outgrowth. To capture the underlying spatiotemporal dynamics, we evaluate both spatial and sequence-based learning approaches. We consider convolutional neural networks (EfficientNet, ResNet, DenseNet, MobileNet, VGG, AlexNet), segmentation-based models (U-Net, U-Net++, Attention U-Net, DeepLabV3, HRNet) and transformer-based architectures (Vision Transformer, Swin Transformer, SegFormer). We investigate transfer learning using both fixed and fine-tuned strategies. Experimental results show that classification performance is greatly enhanced through the integration of temporal information. EfficientNet-B7, EfficientNet-B6, DenseNet-169, and DenseNet201 are consistently better than competing architectures for all tested models. EfficientNet-B7 with the use of temporal sequences input reaches an accuracy of 98.86% with a ROC-AUC of 0.998. The results highlight the significance of spatio-temporal feature learning and the value of DL frameworks in automated classification of dormant versus proliferative breast cancer cells in physiologically relevant microenvironments.
De, S.
Show abstract
Cervical cancer represents a pressing global health challenge, emphasizing the critical need for accurate and timely diagnostic methods to facilitate effective treatment and improve survival rates. In response to this challenge, the study presents CerViX-Net, an innovative classification framework designed to advance cervical cancer detection through enhanced computational efficiency and diagnostic accuracy. The development of CerViX-Net is motivated by the limitations of traditional diagnostic models, particularly in handling the computational and memory demands of large-scale data, while ensuring precise feature extraction and classification. CerViX-Net employs a hybrid deep learning architecture that combines the capabilities of ResNet50, EfficientNet-B0, and a Modified Vision Transformer (ViT) module. The ResNet50 branch extracts hierarchical features through stacked convolutional and identity blocks. In another path, the modified ViT module transforms image patches via linear projection, augments them with positional and class embeddings, and processes them using Parallel Transformer Encoder layers to model contextual relationships. Concurrently, EfficientNet-B0 utilizes MBConv blocks to extract multi-scale representations. The feature outputs from all three branches are integrated and passed through a classification head consisting of dropout layers and dense layers to ensure robust and accurate predictions. The proposed framework is rigorously evaluated on the Mendeley LBC dataset, achieving exceptional performance metrics with an accuracy of 99.69%, precision of 99.28%, recall of 99.48%, and an F1-score of 99.52%. The robustness of CerViX-Net is further validated on the SIPaKMeD and Herlev Pap Smear datasets, where it demonstrates comparable excellence, underscoring its efficacy and adaptability across diverse cytology datasets. Statistical validation using Friedman's test further reinforces its superiority over competing methods.
Singhvi, S.; Singhvi, R.
Show abstract
Medical imaging pipelines routinely copy single-channel grayscale data into three identical RGB channels before classification, usually without justification. This study tests whether that step affects model predictions. Four coordinated experiments on bit-identical RGB inputs sorted eleven classical machine learning models into three groups: five that were invariant to the copy, two that were nearly invariant, and four whose predictions changed. On the Kaggle Alzheimer MRI Dataset (6,400 images, four classes, five seeds), five models (AdaBoost, HistGradientBoosting, KNN, SVM_Polynomial, and SVM_RBF) produced identical predictions in both conditions for every seed, where KNN is k-nearest neighbors and SVM a support vector machine, with polynomial and radial basis function (RBF) kernels. Two models (GaussianNB and SVM_Linear) differed by at most one of 1,280 samples, a dataset-dependent gap rather than exact invariance. The remaining four (DecisionTree, ExtraTrees, RandomForest, and LogisticRegression) differed substantively. A regularization sweep on Logistic Regression traced its gap to a single cause. As L2 regularization weakened, the color-minus-grayscale macro F1 gap shrank steadily, from +12.07 percentage points at C=0.001 to near zero at C=100 (paired Wilcoxon p=0.0020 under strong regularization), showing the effect scales with feature count rather than image content. A replication on the OASIS dataset, matched in size and class balance, reproduced every grouping, and the Logistic Regression gap reappeared in the same direction at smaller magnitude (+5.30 points macro F1). Two deep networks, ResNet18 and DenseNet121, gave identical predictions across all twenty paired conditions. Channel triplication left most models unchanged while multiplying classical training time 2.3 to 4.0 times without benefit.
Gao, Y.; Cui, Y.
Show abstract
Large-scale clinical and biomedical datasets increasingly contain both diverse subgroup attributes (e.g., demographic or clinical subgroups) and multiple prediction targets. Although various machine learning approaches can address subgroup differences or multi-target prediction, they often consider these aspects independently rather than jointly. To more effectively capture the shared and subgroup-specific information in such complex datasets, we propose the Integrative Transfer Network (ITN), a deep neural network designed to leverage data across subgroups and multiple related outcomes simultaneously. In extensive experiments, including time-to-event and classification tasks where demographic subgroups and multiple disease end-points are prevalent, ITN demonstrates consistent improvements in subgroup-specific prediction by borrowing strength from other subgroups and outcomes. We envision ITN as a unified frame-work for learning from heterogeneous datasets where subgroup-specific insights are critical.
Chen, Y.; Yi, H.; Rao, S.; Weber, A.; Hassmiller-Lich, K.; Sylvia, S.
Show abstract
Inappropriate antibiotic use presents a major global health challenge, particularly in low-resource settings where access to quality care is limited but antibiotics remain relatively unrestricted. This study estimates the causal effect of frontline primary care quality on inappropriate community antibiotic use, combining detailed community-based data from approximately 100 rural villages in rural China with an instrumental variable (IV) approach embedded within a double/debiased machine learning (DML) framework. We linked objective measures of village doctor clinical practice quality, measured through unannounced standardized patient visits, to household-level antibiotic use data collected from the same villages. To identify the causal effect, we constructed multiple candidate instruments from extensive provider characteristics and used an ensemble of machine learning algorithms within a flexible DML-IV framework to approximate an optimal instrument, addressing a many-weak-instruments problem. We found that improving village provider clinical practice quality reduced both antibiotic receipt during healthcare encounters for common diseases and household antibiotic storage for future self-medication. Our findings suggest that strengthening frontline primary care quality can meaningfully reduce inappropriate community antibiotic use without restricting access to essential treatment. More broadly, this study illustrates how causal machine learning can strengthen conventional causal estimation in complex observational settings in global health economics research.
Li, W.; Chang, S.; Zhu, L.; Bao, Y.; Liu, T.; Wang, H.; Lin, G. N.
Show abstract
Ground reaction force (GRF)-based gait analysis provides objective, non-invasive evidence for neurological and musculoskeletal assessment, but its translation into medical AI decision support is limited by heterogeneous sensing devices, variable-length recordings, acquisition noise, sensor failures, and restricted access to high-cost gait laboratories. We propose SubGaitNet, a decision-oriented and interpretable AI framework designed to address four clinically relevant challenges in GRF-based medical AI: signal-length variability, sensing noise, long-range gait-phase dependency, and pathological frame-to-frame variability. SubGaitNet integrates GRF temporal slicing, multi-scale deep residual shrinkage, masked Transformer modeling, and a Sub-LSTM branch for adjacent-frame variability modeling. In subject-independent evaluation on two public clinical gait datasets, SubGaitNet achieved an AUC of 0.979 for Parkinson's disease (PD) screening and an ACC of 0.940/F1-score of 0.910 for Hoehn & Yahr severity assessment using wearable pressure insoles. On the GaitRec force-plate dataset, SubGaitNet achieved ACC values of 0.951 and 0.918 for four-class and five-class musculoskeletal impairment assessment, respectively. Additional analyses showed stable bootstrap confidence intervals, calibrated PD screening probabilities (Brier score = 0.059; expected calibration error = 0.051), positive decision-curve net benefit across clinically relevant thresholds, and ordinally plausible H&Y errors. Robustness tests under simulated sensor failure, noise perturbation, and reduced-channel inputs supported the model's stability under clinically plausible sensing uncertainty and accessibility constraints. SHAP explanations highlighted biomechanically meaningful hindfoot and forefoot regions. Overall, SubGaitNet provides a reusable, interpretable, and decision-support-oriented AI methodology for GRF-based gait health assessment, while prospective clinician-in-the-loop validation remains necessary before clinical deployment.
Patel, I.; Leyva, A.; Niazi, M. K. K.
Show abstract
FeePredict is a three-stage random forest machine learning framework to simul-taneously predict whether Medicare reimbursement rates for specific procedures will change, in which direction they will change, and by how much. FeePredict was ap-plied to the four major Medicare fee schedules: the Clinical Laboratory Fee Schedule (CLFS), the Physician Fee Schedule (PFS), the Ambulance Fee Schedule (AFS), and the Durable Medical Equipment, Prosthetics, Orthotics, and Supplies (DMEPOS) fee schedule. Each of these fee schedules contains publicly available data from the Centers for Medicare & Medicaid Services (CMS) for the years 2024, 2025, and 2026, with the number of procedures represented in the data ranging from 3,264 to 2,952,842 observations.FeePredict utilizes lag-1 feature engineering and train-only preprocessing steps to ensure that there is no data leakage into the model. Chronological out-of-time valida-tion was performed on three of the four fee schedules to determine the generalizability of the model over time. FeePredict significantly outperformed the assumption that there would be no changes to Medicare reimbursement rates for procedures (p < 0.001), achieving concordance indices between 0.815 and 0.998, and reducing the mean abso-lute error for predicting changes to reimbursement rates by 29% to 85%. Permutation testing of the model with shuffled reimbursement rate labels indi-cates that there is no evidence of data leakage (AUC values: 0.467-0.515). The model achieved concordance indices of 0.854 and 0.972 for the CLFS and DMEPOS fee sched-ules, respectively, outside of its training period, but performed less well outside of its training period for the PFS, indicating that it generalizes less well to changes to the Medicare policy regime that existed after its training period. Overall, though, these re-sults indicate that it is possible to accurately predict whether Medicare reimbursement rates for medical procedures will change using only data from the historical versions of those fee schedules.
Mahtabi, B.; Nasr-Esfahani, E.; Yaraghi, S.
Show abstract
Pneumonia is a leading cause of infectious disease mortality worldwide, accounting for approximately 2.5 million deaths annually and 15% of deaths in children under five. Chest X-ray imaging remains the primary diagnostic tool, but accurate interpretation requires radiological expertise that is disproportionately concentrated in high-income settings, creating a diagnostic gap where disease burden is highest. Automated deep learning offers a scalable complement to specialist-dependent diagnosis, yet clinical adoption requires both high accuracy and transparent, interpretable reasoning. Convolutional neural networks (CNNs) have shown strong potential for pneumonia detection from chest X-rays, but two barriers impede clinical translation: the interpretability of black-box models and the computational feasibility of large architectures in resource-constrained settings. Explainable AI (XAI) methods such as Grad-CAM, Grad-CAM++, and Score-CAM address the interpretability barrier, yet systematic quantitative comparisons across multiple CNN architectures remain scarce. Furthermore, CNN architectures widely used for medical image classification carry high parameter counts that limit feasibility in resource-constrained settings, motivating architectures that achieve competitive accuracy with substantially fewer parameters. Here we propose a parameter-efficient deep learning framework for pneumonia detection based on transfer learning, evaluated across three CNN architectures representing distinct architectural families: EfficientNet-B0 with fine-tuning (proposed method), ResNet50, and DenseNet121, trained under identical conditions on the Kaggle chest X-ray dataset (5,863 images). Our method achieved 90% classification accuracy, outperforming both baselines while requiring 4.8x fewer parameters than ResNet50. To evaluate explainability, Grad-CAM, Grad-CAM++, and Score-CAM were applied across all three architectures and compared quantitatively using Intersection over Union against manually annotated lung segmentation masks, Insertion score, and Deletion score, with pairwise statistical validation via Wilcoxon signed-rank tests and Bonferroni correction. Findings show that classification accuracy and XAI explanation quality must be evaluated independently, and that the proposed parameter-efficient architecture offers a favorable trade-off for resource-constrained clinical deployment.
Garcia, N. M.
Show abstract
Conventional electrocardiography is highly effective for waveform and rhythm diagnosis, but it is less suited to showing how the internal shape of hundreds or thousands of consecutive heartbeats changes over time. We introduce FOXTAIL, a complementary view that represents each cardiac cycle as an ordered sequence of changes in signal direction. Overlaying these sequences in a fixed visual field makes beat-to-beat organization visible and allows the density, size, stability, and scale persistence of those changes to be measured. We evaluated the representation in recordings containing normal sinus rhythm, paroxysmal atrial fibrillation, severe heart failure, ventricular tachyarrhythmia, and controlled electrode-motion noise. Paired recordings showed that FOXTAIL descriptors can reveal within-person state changes that are not conveyed by a single average beat. The noise and pre-fibrillation analyses also showed that a dense event pattern is not automatically equivalent to physiological complexity, measurement artifact, or impending disease. FOXTAIL is therefore not proposed as a replacement for the diagnostic ECG or as a new classifier, but as an observation and measurement domain for asking a more basic question: how is the electrical organization of the heart changing from one beat to the next, and which of those changes persist across scale?
Komolafe, O. O.; Roberts, A. C.; Shelley, J.; Tawiah, A. K.
Show abstract
High-quality, domain-specific datasets are foundational to advancing educational tools and AI systems in healthcare, yet assembling case repositories from real-world clinical records faces substantial privacy, ethical, and licensing barriers. Synthetic data generation offers a compelling pathway forward, but educational cases require rigorous validation to ensure clinical plausibility and pedagogical utility. This pilot study introduces PhysiCase, a dual-layer validation pipeline for synthetic case generation and evaluates the feasibility of combining automated LLM-based screening with expert educator review. We generated 128 synthetic musculoskeletal(MSK) cases using four frontier large language models (GPT-4.1, GPT-4o, Google Gemini 2.5 Pro, and Llama 4 Scout) across 28 clinical conditions. Cases underwent automated quality screening using an "LLM-as-judge" framework (DeepEval) assessing prompt alignment, JSON correctness, answer relevance, bias, toxicity, and completeness. Ninety cases (70.3%) passed automated filtering and proceeded to expert evaluation by four MSK physiotherapy educators, who rated medical accuracy, realism, fidelity, relevance, and usability on 5-point Likert scales. GPT-4.1 demonstrated the highest automated pass rate (96\%) and strongest expert ratings (medical accuracy 4.10/5, usability 4.38/5), while Llama 4 Scout showed the lowest pass rate (33.3%) and expert ratings. Expert-evaluated cases achieved strong content validity indices for usability (97.5%), relevance (97.5%), and realism (95%), though medical accuracy showed greater variance (CVI 87.5%). Cross-layer correlation analysis revealed that automated completeness metrics moderately aligned with expert usability ratings , while answer relevance and prompt alignment showed weak or negative correlations with clinical correctness. Qualitative analysis identified three primary failure modes: reductive logic, biomechanical inconsistency, and administrative/contextual gaps. The dual-layer validation framework proved methodologically viable: automated screening efficiently reduced expert review burden, while human judgment remained indispensable for detecting subtle clinical reasoning failures. LLM-generated synthetic cases has the potential to meet practical educational needs for MSK physiotherapy, but expert validation is essential to safeguard clinical accuracy. These findings support a scalable division of labour for synthetic case development, with targeted improvements to prompting and automated reasoning checks needed to address identified "nuance gaps." The code for this paper is available on https://github.com/kwid-ai/PhysiCase
Lo, H. U.; Gao, Z.; Loi, H. F.; Cheng, S. K.
Show abstract
Surface electromyography (sEMG) is the most practical non-invasive interface for myoelectric prostheses, exoskeletons, and rehabilitation systems, but power-line interference (PLI) contamination and excessive digital pipeline group delay still limit its clinical adoption. This paper proposes a co-designed analog-digital correction system combining a high-CMRR front-end with an exponentially-windowed RMS (EMRMS) envelope estimator and a recursive single-tone PLI canceller. We present a closed-form CMRR model capturing the electrode-skin imbalance, and provide a complete stability analysis of the LMS canceller. The EMRMS estimator reduces the computational overhead from[O] (L) to strictly[O] (1) in both time and space complexities. Featuring no data-dependent branching, the algorithm achieves deterministic algorithmic execution time (zero jitter under an RTOS environment) and is natively compatible with fixed-point arithmetic on microcontrollers lacking a hardware Floating-Point Unit (FPU). A reference implementation reaches an 8.2 {micro}s median per-sample latency, yielding an end-to-end delay of[~] 30 ms -- leaving a generous >90 ms budget for electromechanical actuation -- while requiring an active CPU duty cycle of merely 1.6%, enabling prolonged deep-sleep intervals. Validation on the public Ninapro DB2 dataset demonstrates a 13.9 dB mean SNR improvement (averaged across 12 channels; single-channel comparison: 9.7 dB, Table 3) and a 70.0 {micro}V envelope RMSE against a length-200 rectangular reference. Paired Wilcoxon signed-rank tests confirm statistical significance (p < 0.001) over static baselines, and Pearson correlation analysis ({rho} = 0.993 {+/-} 0.0002) confirms strict morphological fidelity. The full open-source codebase and benchmarks are publicly released. O_TBL View this table: org.highwire.dtl.DTLVardef@299dc5org.highwire.dtl.DTLVardef@3519a0org.highwire.dtl.DTLVardef@2586aborg.highwire.dtl.DTLVardef@1ac5610org.highwire.dtl.DTLVardef@1465c46_HPS_FORMAT_FIGEXP M_TBL O_FLOATNOTable 3:C_FLOATNO O_TABLECAPTIONQuantitative comparison on a common 60 s segment of Ninapro-like synthetic sEMG (single channel) with a 3 mV 50.3 Hz mains tone slightly drifted from the static notchs design centre at 50.0 Hz, stress-testing the adaptive corrector under a frequency mismatch. The Ninapro multi-channel aggregate (13.9 dB) reported in Section 3.4 uses mains exactly at 50 Hz (matched notch) and so achieves a higher {Delta} SNR. "MAC/sample" excludes the EMRMS square root and the pre-computed LMS sine/cosine. C_TABLECAPTION C_TBL